Ijraset Journal For Research in Applied Science and Engineering Technology
Authors: Amol Chaudhary, Mahika Garg, Roopali Garg
DOI Link: https://doi.org/10.22214/ijraset.2026.84069
Certificate: View Certificate
The accurate estimation of medical insurance expenses is a key requirement for risk assessment, premium determination and resource planning in health insurance systems. Prior research is usually limited by small datasets (986–2,773 records with 6–11 features) and often lacks important health, lifestyle, socioeconomic and behavioural variables, which limits its practical relevance. This study addresses these limitations. A machine learning framework is applied to a heterogeneous dataset comprising 25,000 records and 24 variables spanning demographics, medical history, behavior, economic status, and activity patterns. Missing values and categorical encoding are handled during the preprocessing of the dataset and IQRbased normalisation is applied for handling outliers and scaling. Feature selection and prediction are achieved by employing various learning paradigms such as XGBoost, Random Forest, Gradient Boosting, traditional regression models, deep learning architectures (backpropagation neural network and LSTM) and Conditional Gaussian Bayesian Networks. Model interpret ability is enhanced through explainable AI techniques. We evaluate the performance using R² Score, MAE and RMSE. Comparative analysis shows that larger and more diverse datasets, combined with advanced machine learning techniques, lead to more accurate and realistic cost predictions that provide practical value to insurers and policy makers. This initiative is consistent with SDG 3 (Good health and well-being) and SDG 10 (Reduced inequalities) by facilitating data-driven equitable pricing of health insurance and enhanced access to healthcare funding.
Accurate prediction of medical insurance costs is essential for premium pricing, risk assessment, and policy formulation. Existing studies are often limited by small datasets (fewer than 3,000 records) with only a few basic features such as age, BMI, smoking status, and dependents. These limited datasets fail to capture important real-world factors like lifestyle, socioeconomic conditions, health trends, and high-risk activities, reducing prediction accuracy, fairness, and practical applicability.
To overcome these limitations, this study proposes a scalable machine learning framework using a large dataset of 25,000 patient records with 24 features, covering demographic, clinical, lifestyle, socioeconomic, and insurance-related variables. The study incorporates a comprehensive preprocessing pipeline, advanced regression algorithms, probabilistic feature selection using Conditional Gaussian Bayesian Networks (CGBN), and SHAP (SHapley Additive exPlanations) for model interpretability. Model performance is evaluated using R² Score, Mean Absolute Error (MAE), and Root Mean Squared Error (RMSE).
Previous research on medical insurance cost prediction primarily employed regression and ensemble learning techniques such as Random Forest, Gradient Boosting, and XGBoost on small public datasets with limited features. Explainable AI methods like SHAP and ICE (Individual Conditional Expectation) have been used to identify important predictors such as age, BMI, and smoking status.
Recent studies have integrated explainable AI with ensemble learning to improve prediction accuracy, while hybrid approaches combining Conditional Gaussian Bayesian Networks (CGBN) with machine learning have demonstrated effective feature selection by eliminating redundant variables. Comparative analyses have consistently shown Gradient Boosting to outperform other regression methods in terms of MAE and RMSE.
Despite these advances, existing studies remain constrained by limited dataset size, low feature diversity, and insufficient external validation. This study addresses these shortcomings by integrating a large multidimensional dataset, advanced machine learning algorithms, probabilistic feature selection, and explainable AI within a unified prediction framework.
The study utilizes a Kaggle dataset containing 25,000 anonymized patient records with 24 attributes categorized into:
The dataset provides a comprehensive representation of health, lifestyle, and financial risk factors, making it suitable for realistic insurance cost prediction.
A structured preprocessing pipeline was developed to improve data quality and model reliability:
The study compares multiple predictive models:
To improve feature selection, the study employs Conditional Gaussian Bayesian Networks (CGBN) using algorithms such as HC, Tabu, PC, GS, MMHC, and rsmax2. Model interpretability is enhanced through SHAP, which quantifies each feature's contribution to individual predictions.
Models are evaluated using three standard regression metrics:
The proposed framework consists of two parallel approaches:
The top-performing models from both approaches undergo further analysis using:
This comprehensive analysis improves model transparency and provides valuable insights into factors affecting insurance costs.
Among all regression models, ensemble learning methods demonstrated superior predictive performance.
The results indicate that ensemble learning approaches, such as Random Forest, Gradient Boosting and XGBoost, perform better than standard regression models in terms of accuracy and generalization [1] [6] [7]. We propose an interpretable machine learning framework for medical insurance cost prediction on a large heterogeneous dataset of 25,000 records. To offer high quality inputs for Normal Regression-Based Modelling and a hybrid Conditional Gaussian Bayesian Network (CGBN)-based framework, we developed a single pre-processing pipeline including imputation, encoding, scaling and IQR-based outlier treatment. In the experimental assessment, the ensemble models outperformed the individual models in terms of R2 Score, MAE, and RMSE, obtaining a maximum R2 Score of about 95.64%. Moreover, we established the higher significance of healthcare use and lifestyle-related variables compared to demographic features in predicting insurance cost using SHAP analysis. The hybrid CGBN approach showed how probabilistic feature selection improves interpretability and dimensionality reduction. The prediction accuracy fell in the limited feature space (test R2 ? 77.74%) but remained steady in proportion to the significance of features and improved transparency, indicating the trade-off between accuracy and explainability. The overall results suggest that the scalable and policy relevant medical insurance cost prediction is doable with the combination of robust pre-processing, feature dense datasets, ensemble learning and explainable AI. Future work will formalize a composite Lifestyle and Behavioural Health Risk Index (LBHRI), integrate longitudinal wearable data streams, and incorporate fairness-aware learning to enhance ethical, personalized, and dynamic risk prediction.
[1] U. Orji and E. Ukwandu, “Machine learning for an explainable cost prediction of medical insurance,” Machine Learning with Applications, vol. 15, Art. no. 100516, 2024. doi: 10.1016/j.mlwa.2023.100516. [2] M. M. Billa and T. Nagpal, “Medical insurance price prediction using machine learning,” Journal of Electrical Systems, vol. 20, no. 7, pp. 2270–2279, 2024. [3] A. Rajkomar, J. Dean, and I. Kohane, “Machine learning in medicine,” New England Journal of Medicine, vol. 380, no. 14, pp. 1347–1358, 2019. doi: 10.1056/NEJMra1814259. [4] N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, and A. Galstyan, “A survey on bias and fairness in machine learning,” ACM Computing Surveys, vol. 54, no. 6, pp. 1–35, 2021. doi: 10.1145/3457607. [5] D. Olawade, A. Osborne, A. A. Soladoye, O. E. Oluwadare, E. O. Awogbindin, and O. Z. Wada, “Smart insurance analytics: A novel ensemble feature selection approach to unlock health insurance coverage prediction in Sierra Leone,” International Journal of Medical Informatics, vol. 211, Art. no. 106313, 2026. doi: 10.1016/j.ijmedinf.2026.106313. [6] T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” in Proc. 22nd ACM SIGKDD Int. Conf. Knowledge Discovery and Data Mining, 2016, pp. 785–794. doi: 10.1145/2939672.2939785. [7] R. Genuer and J.-M. Poggi, “Random forests,” in Random Forests with R. Cham, Switzerland: Springer, 2020, pp. 33–55. doi: 10.1007/978-3030-56485-8 3. [8] S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Advances in Neural Information Processing Systems, vol. 30, 2017. [9] P. Biecek and T. Burzykowski, Explanatory Model Analysis: Explore, Explain, and Examine Predictive Models. Boca Raton, FL, USA: Chapman and Hall/CRC, 2021. [10] S. Zou, C. Chu, N. Shen, and J. Ren, “Healthcare cost prediction based on hybrid machine learning algorithms,” Mathematics, vol. 11, no. 23, Art. no. 4778, 2023. doi: 10.3390/math11234778. [11] I. Guyon and A. Elisseeff, “An introduction to variable and feature selection,” Journal of Machine Learning Research, vol. 3, pp. 1157– 1182, 2003. [12] C. J. Willmott and K. Matsuura, “Advantages of the mean absolute error (MAE) over the root mean square error (RMSE) in assessing average model performance,” Climate Research, vol. 30, no. 1, pp. 79–82, 2005. [13] D. Koller and N. Friedman, Probabilistic Graphical Models: Principles and Techniques. Cambridge, MA, USA: MIT Press, 2009. [14] M. Kapse, V. Sharma, R. Vidhale, and V. Vellanki, “Customization of health insurance premiums using machine learning and explainable AI,” International Journal of Information Management Data Insights, vol. 5, no. 1, Art. no. 100328, 2025. doi: 10.1016/j.jjimei.2025.100328. [15] karthikeyanrajuz, “Medical Insurance Prediction Dataset,” Kaggle, [Online]. Available: https://www.kaggle.com/datasets/karthikeyanrajuz/med ical-insurance-prediction-dataset. [Accessed: n.d.]. [16] B. Shickel, P. J. Tighe, A. Bihorac, and P. Rashidi, “Deep EHR: A survey of recent advances in deep learning techniques for electronic health record (EHR) analysis,” IEEE Journal of Biomedical and Health Informatics, vol. 22, no. 5, pp. 1589–1604, 2018. doi: 10.1109/JBHI.2017.2767063. [17] B. Wiley and V. Fuster, “The concept of the polypill in the prevention of cardiovascular disease,” Annals of Global Health, vol. 80, no. 1, pp. 24–34, 2014. doi: 10.1016/j.aogh.2013.12.008. [18] World Health Organization, Global Health Risks: Mortality and Burden of Disease Attributable to Selected Major Risks. Geneva, Switzerland: WHO Press, 2009. [19] Y. Mahwati et al., “Development of machine learning models to predict health insurance claim costs among older Indonesians: A retrospective predictive modeling study,” Journal of Preventive Medicine and Public Health, vol. 59, no. 2, pp. 132–142, 2026. doi: 10.3961/jpmph.25.350. [20] D. K. Kadali et al., “Analysis and prediction of health insurance cost using machine learning approaches,” in Proc. Int. Conf. Computational Innovations and Emerging Trends (ICCIET), 2024, pp. 569–577. doi: 10.2991/978-94-6463-471-6 55.
Copyright © 2026 Amol Chaudhary, Mahika Garg, Roopali Garg. This is an open access article distributed under the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Paper Id : IJRASET84069
Publish Date : 2026-06-30
ISSN : 2321-9653
Publisher Name : IJRASET
DOI Link : Click Here
Submit Paper Online
